Skip to content

test: Add HSTU (Generative Recommenders) Torch AOTI CI test - #8907

Merged
yinggeh merged 4 commits into
mainfrom
yinggeh/tri-1615-add-hstu-ci-test-to-triton-ci
Aug 19, 2026
Merged

yinggeh merged 4 commits into
mainfrom
yinggeh/tri-1615-add-hstu-ci-test-to-triton-ci

Conversation

@yinggeh

@yinggeh yinggeh commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

What does the PR do?

Adds qa/L0_torch_aoti_hstu/test.sh, an end-to-end test for serving an HSTU
(Generative Recommenders) ranking model exported as a Torch AOTI package. It
assembles a platform: "torch_aoti" model repository from the exported package,
starts the FlexKV KV-cache server the model talks to, serves the repository, and
runs the HSTU client against it.

Exporting the package needs the recsys-examples training stack, so it happens in
a separate container in the CI job; the package and the input dump the client
replays arrive here in a volume at /exported_hstu_model. The companion CI job
lands in the internal tritonserver repository.

Checklist

  • PR title reflects the change and is of format <commit_type>: <Title>
  • Changes are described in the pull request.
  • Related issues are referenced.
  • Populated github labels field
  • Added test plan and verified test passes.
  • Verified that the PR passes existing CI.
  • Verified copyright is correct on all changed files.
  • Added succinct git squash message before merging ref.
  • All template sections are filled out.
  • Optional: Additional screenshots for behavior/output changes with before/after.

Commit Type:

  • test

Related PRs:

Internal tritonserver MR adding the L0_torch_aoti_hstu--PyTorch job.

Where should the reviewer start?

qa/L0_torch_aoti_hstu/test.sh is the only file; read it top to bottom. Two
things that look unusual are deliberate: errexit stays off, because the harness
runs tests under bash -ex and the helpers in qa/common/util.sh turn it back
on, which would make the client's exit-code check dead and skip teardown; and
teardown ignores the server's exit status, for the reason under Caveats.

Test plan:

Ran the CI job end to end. Green, with the client reporting max_abs_diff<=0.0625 passed 146/146 batches.

  • CI Pipeline ID: 63551483

Caveats:

  • tritonserver segfaults during shutdown inside an LD_PRELOADed init hook
    shipped by the model's runtime, after the client finishes and after SIGINT.
    Reported to the runtime's owners. Teardown ignores the server's exit status the
    same way kill_server in qa/common/util.sh does.
  • The dataset, the checkpoint and the pinned runtime image the test image builds
    from are all supplied by the CI job.

Background

HSTU serving exercises platform: "torch_aoti" together with a custom backend
runtime and an external KV cache, which no existing L0 test covers.

Related Issues: N/A

Comment thread qa/L0_torch_aoti_hstu/test.sh Outdated
@greptile-apps

greptile-apps Bot commented Jul 29, 2026 •

Copy link
Copy Markdown

Greptile Summary

Adds an end-to-end HSTU Torch AOTI serving test that:

  • Builds a model repository from CI-exported artifacts.
  • Starts and stops the FlexKV cache service.
  • Waits for Triton and its models to become ready.
  • Runs the HSTU client and reports its result.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains; the current startup path waits for Triton's readiness endpoint and handles startup failure before launching the client.

Important Files Changed

Filename Overview
qa/L0_torch_aoti_hstu/test.sh Adds the HSTU test lifecycle, including artifact setup, FlexKV management, Triton readiness handling, client execution, and teardown.

Reviews (12): Last reviewed commit: "Merge branch 'main' into yinggeh/tri-161..." | Re-trigger Greptile

@greptile-apps

greptile-apps Bot commented Aug 7, 2026

Copy link
Copy Markdown

Want your agent to iterate on Greptile's feedback? Try greploops.

L0_torch_aoti_hstu serves the AOTI package that the CI export stages into a
docker volume, using the HSTU runtime baked into the test image. It starts the
FlexKV KV-cache server, loads the exported package as version 1 of
hstu_gr_ranking_kvcache, and replays the dumped inputs through tritonserver,
following Devtech-Compute/distributed-recommender: ci/tritonserver_test.sh.
@yinggeh
yinggeh force-pushed the yinggeh/tri-1615-add-hstu-ci-test-to-triton-ci branch from 1ec12d3 to f03c291 Compare August 12, 2026 01:15
@yinggeh yinggeh added the testing Adding or correcting tests (test: PRs) label Aug 12, 2026
@yinggeh yinggeh self-assigned this Aug 12, 2026
@yinggeh
yinggeh requested a review from whoisj August 12, 2026 02:08
whoisj
whoisj previously approved these changes Aug 12, 2026
whoisj
whoisj previously approved these changes Aug 18, 2026
wait_for_server_ready returns as soon as the process is gone and the
failure branch prints the server log, which is where a link error from
the image's LD_PRELOAD lands.
@yinggeh
yinggeh requested a review from whoisj August 19, 2026 20:53
@yinggeh
yinggeh merged commit 0e006fa into main Aug 19, 2026
3 checks passed
@yinggeh
yinggeh deleted the yinggeh/tri-1615-add-hstu-ci-test-to-triton-ci branch August 19, 2026 20:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

testing Adding or correcting tests (test: PRs)

Development

Successfully merging this pull request may close these issues.

2 participants